Papers with 8B policy model

1 papers
Mutual-Taught for Co-adapting Policy and Reward Models (2025.acl-long)

Copied to clipboard

Challenge: Experimental results show that this iterative approach leads to consistent improvements in both the policy model and reward model.
Approach: They propose a method that iteratively improves both the policy model and reward model without requiring additional human annotation.
Outcome: The proposed method improves both the policy model and reward model without human annotation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations